Papers with base model
Copied to clipboard
| Challenge: | Neural models generate the most common and generic responses all the time . Empirical results show that our method can significantly improve the diversity of responses generated by sequence-to-sequence models. |
| Approach: | They propose an iterative training process and ensemble method based on boosting to improve the diversity of responses generated by neural models. |
| Outcome: | Empirical results show that the proposed method significantly improves diversity and relevance of responses generated by all models. |
Copied to clipboard
| Challenge: | Reinforcement learning from human feedback (RLHF) and reward modeling are key to training powerful large language models (LLMs). |
| Approach: | They propose to combine RLHF and reward modeling to boost model selection . they also demonstrate that a small set of benchmarks could be combined to boost the model selection. |
| Outcome: | The results show that the model selection can be improved by up to 14% compared to the most common (default) choice. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have an array of reasoning capabilities but face limitations such as error propagation and hallucination. |
| Approach: | They propose to use a LLAMA-2 13B CHAT model to act as a task router and task solver to offload certain reasoning steps to external tools that are more suited for the task. |
| Outcome: | The proposed model improves by 35.2% and 5.06% over baseline models and strong GPT-3.5 results. |
Copied to clipboard
| Challenge: | Unsupervised word alignments are not always possible in industrial NLP pipelines, where multilingual annotation guidelines are complex and deviate from semantic consistency due to various factors. |
| Approach: | They propose to constrain word alignment models to remain consistent with both source and target annotation guidelines by leveraging posterior regularization and labeled examples. |
| Outcome: | The proposed model improves on the multiATIS++ dataset over AWESoME, and even a small amount of target language annotations can help. |
Copied to clipboard
| Challenge: | Existing language models have been pre-trained on large-scale code corpora and generate decent code snippets. |
| Approach: | They propose a framework that can provide pre-trained language models with the ability to generate code using private libraries. |
| Outcome: | The proposed framework can generate code using private libraries using off-the-shelf language models or pre-trained models on code corpus containing API information. |
Copied to clipboard
| Challenge: | Domain-adaptive pre-training (DAPT) is one approach for enabling LLMs to handle unseen knowledge. |
| Approach: | They propose to disentangle the answering process into three subtasks and evaluate the performance of each subtask. |
| Outcome: | The proposed model resolves the elicitation task that the base model struggled with but does not resolve other subtasks. |
Copied to clipboard
| Challenge: | Existing approaches to extract aspect terms from review sentences are limited due to lack of annotated data. |
| Approach: | They propose to refine conventional self-training to progressive self-teaching to reduce noise . they use a discriminator to filter the noisy pseudo-labels. |
| Outcome: | The proposed model outperforms baseline models and achieves state-of-the-art performance on four SemEval datasets. |
Copied to clipboard
| Challenge: | Using contextualized token embeddings, we can extract features of propaganda from contextualized embeddnings without fine-tuning the large parameters of the base model. |
| Approach: | They propose a method for detecting fine-grained categories of propaganda in text by generating synthetically generated embeddings from pre-trained language models. |
| Outcome: | The proposed method is used in the first shared task in fine-grained propaganda detection at NLP4IF as Team Stalin. |
Copied to clipboard
| Challenge: | Large Language Models exhibit a progressive left-leaning bias, but can also produce behavior that aligns with socioeconomic groups. |
| Approach: | They analyze whether persona prompting can accurately predict individual voting decisions . they find that they can simulate the voting behavior of European Parliament members reasonably well . |
| Outcome: | The proposed model can predict the voting behavior of European Parliament members reasonably well, with a weighted F1 score of approximately 0.793. |
Copied to clipboard
| Challenge: | Adapters perform dialogue act classification and domain-specific slot tagging in the emergency response domain. |
| Approach: | They propose to build a system that performs dialogue act classification and domain-specific slot tagging while being efficient, flexible and robust. |
| Outcome: | The proposed model performs well in the emergency response domain while being efficient, flexible and robust. |
Copied to clipboard
| Challenge: | Existing methods for decoding text using beam search are expensive and require reinforcement learning. |
| Approach: | They propose a method that allows us to reap the full benefits of beam search with no additional computational cost. |
| Outcome: | The proposed method outperforms greedy decoding and beam search on machine translation tasks with minimal computational cost. |
Copied to clipboard
| Challenge: | Long-Context Language Models (LCLMs) can encode entire document collections, offering a strong alternative to retrieval-augmented generation (RAG). |
| Approach: | They propose to use LCLMs to encode documents with context windows of millions of tokens to improve their performance. |
| Outcome: | The proposed training strategies improve long-context performance and their robustness under compression techniques. |
Copied to clipboard
| Challenge: | Text-to-image (T2I) diffusion models are popular for image manipulation, but also for video generation. |
| Approach: | They propose a novel T2I diffusion model based on latent diffusion that extends the base model for various applications. |
| Outcome: | The proposed model achieves high quality and photorealism and is 3 times faster than the base model. |
Copied to clipboard
| Challenge: | Large language models (LLMs) exhibit strong reasoning capabilities but typically require expensive post-training to reach high performance. |
| Approach: | They propose to use token-level Adaptive Routing to steer frozen LLMs toward structured reasoning entirely at inference time. |
| Outcome: | Extensive experiments show that TARo significantly improves reasoning performance by up to +22.4% over base model and +8.4% . |
Copied to clipboard
| Challenge: | Unlike professional Business-to-Consumer (B2C) e-commerce platforms, consumer-to consumer (C2C), is mainly targeting individual sellers. |
| Approach: | They develop an intelligent product listing tool that generates product descriptions using various product attributes such as category, brand, color, condition, etc. |
| Outcome: | The proposed tool outperforms the base model in domain-specific tasks while producing less hallucination. |
Copied to clipboard
| Challenge: | Existing work adopts separate modules for retrieval and generation, which may be suboptimal since the retrieval task and generation task cannot benefit from each other to improve performance. |
| Approach: | They propose a backbone-shared RAG framework that uses a domain-specific corpus to continuously pre-train a model and then trains two plug-and-play Low-Rank Adaptation modules based on the shared backbone to minimize retrieval and generation losses respectively. |
| Outcome: | The proposed framework outperforms baseline models by 5% and 13% in Hit@3 upon two datasets in retrieval evaluation and by 23% in terms of BLEU-3 in generation evaluation. |
Copied to clipboard
| Challenge: | Quantization Aware Training (QAT) is expensive to train and unscalable to large models. |
| Approach: | They propose a parameter-efficient framework targeting per-channel 4-bit weight-activation quantization of large language models. |
| Outcome: | The proposed framework preserves accuracy within 0.11 percentage points of the full-precision baseline on Llama-2-7B zero-shot tasks while training only 1.26% of total parameters. |
Copied to clipboard
| Challenge: | Current ELS’s are not sufficiently effective, possibly introducing unresolved ambiguities and irrelevant entities. |
| Approach: | They propose an off-the-shelf entity linking system to extract linked entities and propose Entity2Topic (E2T) module attachable to a sequence-to-sequence model that transforms a list of entities into a vector representation of the topic of the summary. |
| Outcome: | The proposed model improves the performance of the Gigaword and CNN summarization datasets by at least 2 ROUGE points. |
Copied to clipboard
| Challenge: | Experimental results show that ProConSuL significantly improves code summaries and reduces the number of hallucinations. |
| Approach: | They propose a framework to provide a large language model with precise information about the code structure from program analysis methods. |
| Outcome: | The proposed framework significantly improves code summaries and reduces hallucinations compared to the base model. |
Copied to clipboard
| Challenge: | Experimental results show that our approach significantly outperforms the supervised counterparts, and can even achieve competitive performance to supervised state-of-the-art (SoA) model. |
| Approach: | They propose a syntactic and semantic-driven learning approach that can learn open IE models without human-labelled data by leveraging syntakic and semantic knowledge as noisier, higher-level supervision. |
| Outcome: | The proposed approach outperforms supervised counterparts and can achieve competitive performance to supervised state-of-the-art models. |
Copied to clipboard
| Challenge: | Existing post-training pipelines that generate QA pairs require costly expert annotation and synthetic data that drops evidence structure. |
| Approach: | They propose a system that converts raw biomedical papers into evidence-enriched training sets and a domain-specialized VLM. |
| Outcome: | Ryze synthesizes QA pairs with complete supporting evidence, reduces layout and OCR errors . the system outperforms the base model on LAB-Bench and surpasses GPT-5.2 by +3.8%. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have greatly improved the performance on most natural language tasks, and often show surprisingly good zero-shot generalization to new domains. |
| Approach: | They propose to continuously pretrain the Llama 3.1 base models on 1 trillion tokens of e-commerce data to introduce domain specific knowledge into the model while at the same time keeping the general capabilities intact. |
| Outcome: | The proposed model can be adapted to the new domain without sacrificing performance on general domain tasks. |
Copied to clipboard
| Challenge: | Large Vision Language Models lack domain-specific data for reasoning on complex problems. |
| Approach: | They propose to use explicit knowledge-infused questions, answers, and reasons to answer and reason upon the questions. |
| Outcome: | The proposed model improves by 25% over the baseline model. |
Copied to clipboard
| Challenge: | Recent studies have shown that most abstractive summarization models are unfaithful and suffer from a wide range of hallucination. |
| Approach: | They propose a candidate summary generation and ranking technique to improve summary factuality without sacrificing quality. |
| Outcome: | The proposed method shows that the model trained using the proposed method improves on factuality and similarity-based metrics without conflicting with the model. |
Copied to clipboard
| Challenge: | a finetuned model may be better base models than the vanilla pretrained model . this scheme, often referred to as intertraining, is the focus of the present work . |
| Approach: | They propose a scheme to analyze the potential intertraining gain independently for the target dataset and for a base model being considered as a starting point. |
| Outcome: | The proposed model is strong even if training data was not aligned with target dataset. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are rapidly being adopted across various domains, but adoption in the regulated banking industry is limited due to their tendency to hallucinate, exhibit over-agreeable behavior, and lack alignment with domain-specific knowledge and constraints. |
| Approach: | They propose a framework for training grounded domain-specific LLMs that optimizes answer quality, citation grounding, and calibrated refusal under real-world deployment constraints. |
| Outcome: | The proposed model outperforms GPT-4.1 on citation grounding and calibrated refusal under real-world deployment constraints. |
Copied to clipboard
| Challenge: | Large-scale language models (LLMs) like ChatGPT have demonstrated impressive abilities in generating responses based on human instructions. however, their use in the medical domain can be challenging due to their lack of specific, in-depth knowledge. |
| Approach: | They propose a system that integrates authoritative medical textbooks into LLMs’ framework using plug-and-play modules. |
| Outcome: | The proposed system outperforms the specialized Med-PaLM 2 model on three medical QA tasks by 11.6% to 16.6%. |
Copied to clipboard
| Challenge: | Existing frameworks for multilingual modeling face communication costs and parameter interference conflicts. |
| Approach: | They propose a communication-efficient federated learning framework with low-rank adaptation and language family clustering for Multilingual Modeling (MM) they maintain the weights of the base model, updating the lightweight Low-rank adapt parameters to minimize communication costs. |
| Outcome: | The proposed model outperforms the baseline models in performance and reduces communication overhead. |
Copied to clipboard
| Challenge: | Large language models exhibit significant performance discrepancies between high- and low-resource languages. |
| Approach: | They present an open-source multilingual LLM with 8 billion parameters and a multilingual instruction dataset. |
| Outcome: | The proposed model achieves consistent multilingual representations across languages. |
Copied to clipboard
| Challenge: | Recent research shows that Transformer-style models can be made more efficient by sharing parameters over blocks. |
| Approach: | They propose a framework for distilling latent reasoning into a multiscale jump model that enables flexible test-time compute. |
| Outcome: | Experiments on ARC-AGI show that the proposed model achieves competitive accuracy compared to recursive baselines while requiring fewer sequential updates. |
Copied to clipboard
| Challenge: | Existing research focuses on enhancing large language models through scaling laws or fine-tuning strategies, but ignores the potential of using agent paradigms to compensate for the inherent weaknesses of small models. |
| Approach: | They propose to use structured agent frameworks to improve effectiveness over direct prompting . they also propose to employ routing-based multi-agent systems with collaborative capabilities . |
| Outcome: | The proposed model significantly outperforms direct prompting with single-agent systems . the proposed model is more reliable and cost-effective than other models . |
Copied to clipboard
| Challenge: | Recent advances in NLP have led to the use of pre-trained Transformer models for transfer learning tasks becoming the most common way to solve target tasks. |
| Approach: | They propose a 3-phase technique to adjust a base model for a classification task by adapting the model’s signal to the data distribution and a new data augmentation approach for Supervised Contrastive Learning to correct the unbalanced datasets. |
| Outcome: | The proposed method is compared with other methods and compares it with other approaches. |
Copied to clipboard
| Challenge: | Existing approaches to handle wrong labeling and long-tail relations are labor-intensive and scarce training data. |
| Approach: | They propose a neural network to handle wrong labeling and long-tail relations by collaborating relation-augmented attention. |
| Outcome: | The proposed neural network improves the state-of-the-art on the NYT dataset . |
Copied to clipboard
| Challenge: | Large language model (LLM) based search agents are more likely to produce harmful outputs than base models. |
| Approach: | They propose a query-level shaping term that rewards safe queries and penalizes unsafe ones. |
| Outcome: | The proposed approach reduces harmfulness by over 70% across three red-teaming datasets while producing safe, helpful responses. |
Copied to clipboard
| Challenge: | Large Language Model (LLM) agents finetuned with supervised finetuning may over-commit towards seemingly plausible but suboptimal actions due to limited action space exploration. |
| Approach: | They propose a self-taught actioN deliberation framework that allows LLM agents to explicitly deliberate over candidate actions before committing to one. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on two representative interactive agent tasks and achieves an average 20% improvement over initial finetuning. |
Copied to clipboard
| Challenge: | Conditional random fields (CRF) for label decoding have been a problem for many tasks. |
| Approach: | They propose a two-stage label decoding framework that model long-term label dependencies while being much more computationally efficient. |
| Outcome: | The proposed method outperforms the CRF-based methods and greatly accelerates the inference process. |
Copied to clipboard
| Challenge: | Large language models (LLMs) generate plausible-sounding responses that are factually incorrect. |
| Approach: | They propose an approach to learn more reliable reward models by modifying how unfamiliar finetuning examples are supervised to influence model responses to unfamiliar queries. |
| Outcome: | The proposed approach improves the efficacy of RL factuality finetuning in long-form biography and book/movie plot generation tasks. |
Copied to clipboard
| Challenge: | Existing methods for dependency parsing are often of the pseudo-annotation type, but they fail to consider the change of model structure for domain adaptation. |
| Approach: | They propose a method that accomplishes unsupervised cross-domain dependency parsing without using labeled data. |
| Outcome: | The proposed method achieves consistent performance improvement on CODT1 and CTB9 domains. |
Copied to clipboard
| Challenge: | coding tasks require generated code to be fully executable and functionally correct . current agentic approaches struggle with multi-stage planning, generating, and debugging . |
| Approach: | They propose a framework for LLM agents to efficiently explore the search space in different stages of the code generation process. |
| Outcome: | The proposed framework achieves top results on 7 code generation benchmarks and a 31.9% solving rate on the SWEBench benchmark. |
Copied to clipboard
| Challenge: | a recent study shows that fine-tuning improves the performance of language models . large language models generate acceptable texts in a number of scenarios, a study shows . |
| Approach: | They show that fine-tuning improves the task of hate speech counter-narrative generation . they provide a subset of arguments and a good base model is required for the fine-uning to have a positive impact. |
| Outcome: | The proposed model produces counter-narratives that are as satisfactory as the whole set. |
Copied to clipboard
| Challenge: | Tokenization is a foundational step for Large Language Models (LLMs) but low compression rate of vanilla tokenizers decelerates training and inference process. |
| Approach: | They propose a method to replace the vocabulary of Large Language Models (LLMs) by learning a one-to-one mapping matrix for token IDs. |
| Outcome: | The proposed method significantly improves multilingual text compression rates and vocabulary initialization for Large Language Models. |
Copied to clipboard
| Challenge: | Existing studies show that supervised training is still necessary for complex reasoning tasks. |
| Approach: | They propose a method to integrate uncertainty-based active learning and LoRA to effectively integrate the two methods. |
| Outcome: | The proposed approach outperforms baseline models on three reasoning tasks. |
Copied to clipboard
| Challenge: | Existing methods for domain-specific reasoning with large language models require updating parameter updates. |
| Approach: | They propose a plug-and-play intervention framework that adaptively steers LLM reasoning in activation space. |
| Outcome: | The proposed framework achieves zero-shot accuracy improvements of 3.4–6.5% over the base model while outperforming chain-of-thought-style reasoning with 2–3 higher token efficiency and robust accuracy gains. |
Copied to clipboard
| Challenge: | Existing algorithms for post-training large datasets are requiring a large computational effort. |
| Approach: | They propose to model the changes at logits level during post-training using a separate neural network . they demonstrate that the value network can be seamlessly integrated with another pre-trained model . |
| Outcome: | The proposed model can be integrated with another pre-trained model during inference, enabling similar capability enhancements. |
Copied to clipboard
| Challenge: | Reinforcement Learning methods for text-based games fail to generalize on unseen games, especially in small data regimes. |
| Approach: | They propose a Context Relevant Episodic State Truncation method for irrelevant token removal in observation text for improved generalization. |
| Outcome: | The proposed method shows that it can generalize on unseen games using 10x-20x fewer training games compared to previous state-of-the-art methods despite requiring fewer number of training episodes. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have achieved remarkable success across diverse tasks through large-scale pretraining. |
| Approach: | They propose a framework that filters noisy components from LoRA updates via subspace similarity with the base model. |
| Outcome: | The proposed framework improves accuracy by 12%, reduces forgetting by 29%, and filters out over 30% of LoRA parameters identified as noisy. |
Copied to clipboard
| Challenge: | Continued pretraining (CPT) is a practical route to language adaptation, but improvements on demanding capabilities such as mathematical reasoning are limited. |
| Approach: | They propose to use CPT to adapt large language models to African languages . they use math, code, and synthetic translated data to analyze their models . |
| Outcome: | The proposed models improve on multilingual benchmarks and document-level translation. |
Copied to clipboard
| Challenge: | Recent results show that the mix-of-experts architecture is parameter inefficient . large-scale pre-trained language models can achieve excellent performance in many NLP tasks. |
| Approach: | They propose to build a parameter-efficient mix-of-experts architecture by sharing information across experts. |
| Outcome: | The proposed architecture increases model capacity without increasing computation costs. |
Copied to clipboard
| Challenge: | Existing approaches to cross-lingual summarization use only a single reference, resulting in an underrepresented hypothesis space. |
| Approach: | They propose to use pseudo-labels to regularize cross-lingual summarization training by combining a single reference and a network to perform the model training. |
| Outcome: | The proposed approach significantly improves over gold reference training in XLS with 8 languages from different families. |
Copied to clipboard
| Challenge: | prevailing taxonomies neglect robustness and honesty, yielding safer-on-paper but less useful systems. |
| Approach: | They propose a soft-gating pipeline where a guardian predicts a binary risk label plus a concise explanation and prepends this advice to the original query for re-inference. |
| Outcome: | The proposed model maintains safety while reducing over-refusal. |
Copied to clipboard
| Challenge: | Sequence-to-sequence neural networks have enabled great progress in abstractive summarization. |
| Approach: | They propose to train a second-stage model performing re-ranking on a set of summary candidates by using a mixture of experts. |
| Outcome: | The proposed model outperforms the base model on CNN- DailyMail, XSum and Reddit TIFU with a base PEGASUS. |
Copied to clipboard
| Challenge: | Existing bilingual or multilingual medical LLMs are limited in multilingual data and therefore perform poorly in non-English languages such as Japanese and Chinese. |
| Approach: | They propose to use a trilingual (English, Japanese, Chinese) large language model adapted for the bio-medical domain to harness the knowledge and abilities of the base model. |
| Outcome: | The proposed model can support English, Japanese, and Chinese and is adapted for a bio-medical domain. |
Copied to clipboard
| Challenge: | Existing approaches to align Large Language Models with human preferences are complex and unstable. |
| Approach: | They propose a new approach that maximizes the marginal log-likelihood of a preferred text output by using the preference pair as samples for approximation. |
| Outcome: | The proposed approach maximizes the marginal log-likelihood of a preferred text output, using the preference pair as samples for approximation, and forgoes the need for both an explicit reward model and entropy maximization. |
Copied to clipboard
| Challenge: | Existing methods for interpreting, augmenting, and querying semi-structured tables require pretraining on tables or special model architecture design. |
| Approach: | They construct a dataset with a variety of tables and tasks for instruction tuning and evaluating LLMs. |
| Outcome: | The proposed model achieves comparable or better performance on 7 out of 8 in-domain tasks compared with the base model on 6 out-of-domain datasets. |
Copied to clipboard
| Challenge: | Extensive experiments on challenging mathematical reasoning benchmarks demonstrate that these human-inspired strategies synergistically and significantly enhance performance. |
| Approach: | They propose to use Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation to improve model performance. |
| Outcome: | Extensive experiments on mathematical reasoning benchmarks show that the proposed strategies synergistically and significantly improve performance over the baseline model. |
Copied to clipboard
| Challenge: | Existing work focuses on domain-specific enhancements during fine-tuning, the challenge of which lies in catastrophic forgetting of knowledge across other domains. |
| Approach: | They propose a data composition framework that allows LLMs to enhance their multi-domain capabilities during supervised fine-tuning. |
| Outcome: | The proposed framework improves multi-domain fostering performance by 29.77% compared to uniform weights. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) generate content that can be untruthful or harmful. |
| Approach: | They propose a method that leverages model feedback for alignment . they use a base language model to generate initial responses, critiqued and refined . |
| Outcome: | The proposed method outperforms strong baselines across diverse tasks and model sizes. |
Copied to clipboard
| Challenge: | Unlike token-level likelihood search, which is myopic and often rewards verbosity, our approach works at an intermediate granularity. |
| Approach: | They propose a lookahead quality gate for speculative decoding that accepts the longest reliable prefix of each k-token lookaheaded draft. |
| Outcome: | The proposed method improves accuracy over baselines while achieving 2.6-7.9 faster generation on math and science benchmarks. |
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) is a powerful framework for knowledge-intensive tasks, but its effectiveness in long-context scenarios is often bottlenecked by the retriever’s inability to distinguish sparse yet crucial evidence. |
| Approach: | They propose a framework that fine-tunes the retriever for Answer Alignment by identifying high-quality positive chunks by evaluating their sufficiency to generate the correct answer. |
| Outcome: | The proposed framework improves 14.5% over the base model and maintains strong efficiency for long-context RAG. |
Copied to clipboard
| Challenge: | Existing approaches to few-shot Question Generation (QG) are limited and require manual annotation. |
| Approach: | They propose to use multilingual BERT to perform few-shot question generation with cross-lingual transfer. |
| Outcome: | The proposed model improves in few-shot QG and human evaluation confirms it. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown significant progress in information extraction tasks due to lack of labeled data for fine-tuning and unlabeled text for pre-training. |
| Approach: | They propose a framework in which large language models are fine-tuned to use English translations of low-resource language data. |
| Outcome: | The proposed model improves cross-lingual transfer over the base model on 12 multilingual IE datasets spanning 50 languages. |
Copied to clipboard
| Challenge: | a recent study has shown that GPT-3 fine-tuning models with limited examples is effective . a contrastive learning framework clusters inputs from the same class under different augmented “views” and repels those from different classes. |
| Approach: | They propose a supervised contrastive framework that clusters inputs from the same class under different augmented "views" they combine a contrastive loss with the standard masked language modeling loss in prompt-based few-shot learners . |
| Outcome: | The proposed framework improves on the state-of-the-art methods in a diverse set of 15 language tasks. |
Copied to clipboard
| Challenge: | Medical professionals often query over clinical notes to find information that can support their decision making. |
| Approach: | They propose to use expert-annotated question templates and existing i2b2 annotations to create emrQA, the first large-scale dataset for question answering based on clinical notes. |
| Outcome: | The proposed system can answer clinical questions without using domain knowledge. |
Copied to clipboard
| Challenge: | Existing methods to capture unintended dataset biases are expensive and require elaborate balancing strategies. |
| Approach: | They propose a model-agnostic text classification debiasing framework which can effectively avoid employing data manipulations or designing balancing mechanisms. |
| Outcome: | The proposed framework can effectively avoid data manipulations or designing balancing mechanisms to capture unintended dataset biases. |
Copied to clipboard
| Challenge: | GUI agents have demonstrated remarkable progress in automating complex user interface interactions . training such agents for long-horizon tasks remains challenging due to limited rewards and prohibitive costs. |
| Approach: | They propose a method that leverages expert trajectories as environment experiences for on-policy multi-turn training. |
| Outcome: | The proposed method achieves significant gains over the base model with 1K public trajectories as RL experiences . it achieves competitive performance against strong baselines such as UI-TARS-7B and GPT-4o . |
Copied to clipboard
| Challenge: | Existing models are often used as black boxes to adapt to new domains, but there is no single recipe for making them work. |
| Approach: | They propose to use black box models to improve their performance on new domains by leveraging explanations of their behavior. |
| Outcome: | The proposed method improves model generalization performance on two tasks using explanations. |
Copied to clipboard
| Challenge: | Pre-processing tools such as optical character recognition (OCR) can map document image inputs to textual tokens, then large language models (LLMs) can reason over text. |
| Approach: | They propose a method that integrates outputs of OCR tools and larger multimodal models as intermediate "rationales" a student model is trained to predict rationales and answers based on visual documents . |
| Outcome: | The proposed model outperforms the base model on three visual document understanding benchmarks with only 1% higher computational cost. |
Copied to clipboard
| Challenge: | Eligibility criteria (EC) are critical components of clinical trial design, specifying parameters for participant inclusion and exclusion. |
| Approach: | They propose a method that utilizes Retrieval-Augmented Fine-Tuning to generate structured and cohesive EC directly from clinical trial titles and descriptions. |
| Outcome: | The proposed method outperforms Llama-3.1-8B-Instruct and Llm-as-a-Judge models in BERTScore and EC score. |
Copied to clipboard
| Challenge: | Existing methods for developing LLMs are constrained by static data or sparse reward signals in online settings. |
| Approach: | They propose a framework that iteratively refines tutor agents using a multi-horizon reward function within a dynamic teacher-student simulation environment. |
| Outcome: | The proposed framework improves model performance and balances principles and effectiveness compared to baselines. |
Copied to clipboard
| Challenge: | MLLMs lack visual grounding mechanism to read text embedded in images, or rely on parametric shortcuts . despite strong OCR capabilities, models suffer performance degradation of 12.7% in the VQ setting . |
| Approach: | They propose a plug-and-play training strategy that invalidates shortcuts in text prompts . they propose 'vq' setting where text queries are rendered directly onto images . |
| Outcome: | The proposed training strategy surpasses the base model by 5.4% and GRPO based on original images by 2.7% on four representative OOD benchmarks. |
Copied to clipboard
| Challenge: | Existing provenance detection methods for large language models are infeasible for already published models and compare outputs using hand-crafted or random prompts. |
| Approach: | They propose a detection framework that constructs fingerprints by exploiting LLMs’ inherent vulnerability to prompt injection. |
| Outcome: | The proposed framework achieves high true positive rates while keeping false positive rates near zero. |
Copied to clipboard
| Challenge: | Existing approaches to sentence embeddings are based on contrastive learning (CL) . |
| Approach: | They propose a framework which performs contrastive learning under the self-training paradigm with knowledge distillation and propose 'Group-P shuffling strategy' and averaging logits from multiple teacher components. |
| Outcome: | The proposed framework outperforms many strong baseline methods and yields a new state-of-the-art performance. |
Copied to clipboard
| Challenge: | Existing methods for decomposing fine-tuned LLMs are sensitive to the magnitude of delta values. |
| Approach: | They propose a hierarchical quantization framework that shares low-bit integer weights across similar models. |
| Outcome: | The proposed framework achieves an average accuracy degradation of approximately 3% on fine-tuned models across mathematics, coding, chatbot, and Chinese LLMs. |
Copied to clipboard
| Challenge: | Existing models that use text attributes to improve sentiment classification use text as a categorical feature. |
| Approach: | They propose to represent attributes as chunk-wise importance weight matrices and consider four locations to inject attributes. |
| Outcome: | The proposed method outperforms the state-of-the-art and outperformed previous models. |
Copied to clipboard
| Challenge: | Current instruction tuned language models are trained on textual preference data and therefore not aligned to speech domain. |
| Approach: | They propose to use radio-industry best practices to prompt and learn speech-based preference data to improve speech-suitability of popular instruction tuned language models. |
| Outcome: | The proposed methods achieve the best win rates in head-to-head comparisons, resulting in preferred or tied to the base model in 76.2% of comparisons on average. |
Copied to clipboard
| Challenge: | Prior work focused on building multilingual models that cover a broad spectrum of languages. |
| Approach: | They conduct systematic experiments on how design choices impact the adapted LLM, both in terms of efficiency and end task performance. |
| Outcome: | The proposed model performs better on English-centric models than multilingual models despite poor performance on low-resource languages. |
Copied to clipboard
| Challenge: | Existing methods to generate paraphrases are not trivial and often fail in practice. |
| Approach: | They propose to use imitation learning to boost the performance of generating paraphrases by using a pointer-generator model. |
| Outcome: | The proposed model outperforms the state-of-the-art methods on the benchmark datasets. |
Copied to clipboard
| Challenge: | Unified Multimodal Models have achieved remarkable success in cross-modal comprehension, but a gap persists in their ability to translate internal knowledge into faithful and controllable synthesis. |
| Approach: | They propose a self-improvement framework that partitions a single UMM into three collaborative roles: Proposer, Solver, and Judge. |
| Outcome: | The proposed framework improves on TIIF, DPG, CompBench and UniCycle benchmarks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) increase test-time computation, often in the form of chain-of-thought (CoT) however, reasoning traces can become unnecessarily long, increasing computation costs without improving accuracy and sometimes even degrading performance. |
| Approach: | They propose a multi-stage efficient reasoning method that combines supervised fine-tuning with reinforcement learning using an adaptive length penalty. |
| Outcome: | The proposed method reduces response length by an average of 28% for 8B models and 40% for 32B models while incurring only minor performance drops of 1.6 and 2.5 points, respectively. |
Copied to clipboard
| Challenge: | Large language models (LLMs) require alignment to effectively and safely follow user instructions. |
| Approach: | They propose a simple, training-free algorithm that aligns any base model at inference time using a small aligned model. |
| Outcome: | The proposed algorithm outperforms large aligned models on open-instruction tasks without training. |
Copied to clipboard
| Challenge: | Excessive safety can lead to over-refusal, where models reject harmful-looking yet benign queries, severely limiting utility. |
| Approach: | They propose a lightweight training-based approach that reshapes the distributions of harmful and benign samples within the model’s decision space by using a single-token prefix. |
| Outcome: | The proposed approach can distinguish between harmful and benign samples while keeping the model frozen. |
Copied to clipboard
| Challenge: | Several studies claim that domain-adaptive pretraining improves performance on downstream medical tasks. |
| Approach: | They compare medical LLMs and VLMs against their corresponding base models . they find that medical Lms outperform their base models in 12.1% of cases . |
| Outcome: | The proposed models outperform their base models on medical questions and tasks in 12.1% of cases and reach a tie in 49.8% of cases. |
Copied to clipboard
| Challenge: | Existing work suggests that the degree of hallucination depends on factual errors in training data. |
| Approach: | They propose a method to use training data to reduce hallucination by ensembling parameter variations in training data. |
| Outcome: | The proposed method improves on XSUM and CNN/DM datasets on human evaluations and factual metrics. |
Copied to clipboard
| Challenge: | Large language models (LLMs) still exhibit significant deficiencies in basic language understanding and manipulation. |
| Approach: | They propose a bilingual benchmark to assess the performance of Large language models . they use a set of 15 simple text editing tasks to examine their capabilities . |
| Outcome: | The proposed benchmark aims to assess the performance of Large language models in basic language tasks. |
Copied to clipboard
| Challenge: | Existing studies on unsupervised headline generation focus on a standard dataset and mono-style corpora. |
| Approach: | They propose an unsupervised approach for stylistic headline generation using a pretrained BART model decorated with adapters responsible for different styles. |
| Outcome: | The proposed method separates the task of style learning and headline generation, allowing for the generation of diverse headlines with diverse styles. |
Copied to clipboard
| Challenge: | Parameter-efficient fine-tuning (PEFT) is a common method for fine- tuning large language models . however, once updated, PEFT modules suffer performance degradation on newer versions . |
| Approach: | They propose a method that enhances the PEFT module by focusing on the task-specific pattern while reducing its dependence on certain knowledge in the base model. |
| Outcome: | Experiments show that PEFT modules can maintain performance on updated models without re-tuning . the proposed approach can be used in real-world applications with large model sizes . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been successful in machine translation, but lack of high-quality parallel corpora and cost constrain scalability. |
| Approach: | They propose an LLM-driven dual-learning framework that enables autonomous translation . they employ a robust semantic-aware reward function that balances adequacy with reconstruction fidelity . |
| Outcome: | The proposed model outperforms larger models on benchmarks and achieves parity with state-of-the-art supervised baselines on mainstream benchmarks. |
Copied to clipboard
| Challenge: | Reasoning-capable large language models (LLMs) have driven a major shift in artificial intelligence . these models generate long CoTs, capturing reasoning behaviors such as self-reflection, self-correction, and hypothesis testing. |
| Approach: | They propose a sample-efficient, two-stage training strategy to build reasoning LLMs . they "warm up" a model by distilling Long CoTs from a toy domain to acquire general reasoning skills . |
| Outcome: | The proposed training strategy outperforms existing models on a range of tasks. |
Copied to clipboard
| Challenge: | Existing evaluations test factual medical knowledge in isolation or assess patient-level reasoning without verifying correctness, leaving a critical gap. |
| Approach: | They propose a benchmark that links MIMIC-IV EHRs to a unified knowledge base built from UMLS and other biomedical vocabularies. |
| Outcome: | The proposed model improves by +16.4 macro-F1 points over the base model and eliminates truth inversion errors. |
Copied to clipboard
| Challenge: | Existing methods for linguistic style control lack fine-grained control, require extensive computation, or introduce significant latency. |
| Approach: | They propose a parameter-space approach that extracts style-specific representations by analyzing parameter differences between models trained on contrasting styles and incorporates them into a model with precise control over style intensity. |
| Outcome: | The proposed approach achieves three key capabilities while achieving optimal computational efficiency. |
Copied to clipboard
| Challenge: | Recent studies have shown that strong natural language understanding models are prone to relying on unwanted dataset biases without learning the underlying task. |
| Approach: | They propose two learning strategies to train neural models that are more robust to dataset biases and transfer better to out-of-domain datasets. |
| Outcome: | The proposed methods improve robustness in all settings and transfer better to out-of-domain datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are difficult to interpret due to their black-box nature and randomness. |
| Approach: | They propose a new method which enhances influence functions by addressing fitting errors by eliminating knowledge bias present in the base model before fine-tuning. |
| Outcome: | The proposed method outperforms existing methods and achieves an average AUC of 91.64%. |
Copied to clipboard
| Challenge: | Unsupervised cross-lingual transfer is a process of transferring knowledge between languages without explicit supervision. |
| Approach: | They propose a framework that combines lexical and syntactic knowledge to enhance learning . they use a code-switching technique to implicitly teach lexica and a syntaktic-based graph attention network to help encode syntakic structure. |
| Outcome: | The proposed framework outperforms baselines of zero-shot cross-lingual transfer with 1.0 3.7 points on text classification, named entity recognition, and semantic parsing tasks. |
Copied to clipboard
| Challenge: | Document-level event argument extraction is a challenging task for cross-sentence inference . previous work focused on document-level EAE, but recent work focused more on documentlevel . |
| Approach: | They propose a document-level event argument extraction model that captures contextual clues and latent role information. |
| Outcome: | The proposed model outperforms existing methods on two public datasets with 1.13 F1 and 2.64 F1 improvements on RAMS and WikiEvents respectively. |
Copied to clipboard
| Challenge: | Existing attempts to outline generation are limited by response pair requirements and substantial computation costs. |
| Approach: | They propose a token-level preference self-alignment optimization for outline controllable generation that extends the Bradley-Terry model from pair-wise to list-wise comparison. |
| Outcome: | The proposed method outperforms existing methods by 19.28% in performance while requiring only 56.25% training time. |
Copied to clipboard
| Challenge: | Existing evaluation methods for large language models (LLMs) are inadequate to provide solid conclusions for key experiments such as data ablation and scaling law. |
| Approach: | They propose a method specifically designed to optimize the evaluation of base models by incorporating two innovations: In-Context Light-instruction Prompt and Blank-ppl for multi-choice tasks with candidate options. |
| Outcome: | The proposed method significantly improves stability and consistency of evaluations during pre-training and consistency between base and instruct models. |
Copied to clipboard
| Challenge: | Recent work has focused on improving the mathematical reasoning capabilities of Large Language Models (LLMs). |
| Approach: | They propose an end-to-end framework to integrate FL into NL math reasoning . they propose a problem alignment method that reformulates QA and existence problems . |
| Outcome: | The proposed framework achieves 89.80% and 84.34% accuracy rates on the MATH-500 and the AMC benchmarks. |
Copied to clipboard
| Challenge: | emergence of large language models (LLMs) has brought about new opportunities for machine translation. |
| Approach: | They propose a method for data curation that supplements the infrequent senses of polysemous words. |
| Outcome: | The proposed method outperforms established baselines on the WMT2022 test sets and is applicable to other pre-trained models. |
Copied to clipboard
| Challenge: | Existing vision-language pre-training models use multi-modal encoders to encode image and text, causing noisy training corpora. |
| Approach: | They propose a vision-language pre-training framework with two autoencoders for efficient training . they propose masked tokens and a gated interaction mechanism to cope with noise . |
| Outcome: | The proposed model achieves 2.2% R@1 gains on COCO Text Retrieval and 1.1% on refCOCO+ on six datasets. |
Copied to clipboard
| Challenge: | Recent approaches to structured NLP tasks use autoregressive models trained on pairs of unstructured input text and structured output targets. |
| Approach: | They propose a model that combines constrained and unconstrained decoding in two phases to achieve two weak predictions. |
| Outcome: | The proposed model outperforms previous approaches both in and out of distribution, addressing several common errors identified in those approaches. |
Copied to clipboard
| Challenge: | Large language models (LLMs) with their extensive parameters and high memory demands are challenging to fine-tune for specific applications with limited resources. |
| Approach: | They propose a method that dynamically adjusts the adapter’s rank using rank-subspace analysis, optimizing performance with fewer parameters. |
| Outcome: | The proposed method improves model accuracy with minimal parameter changes and demonstrates the importance of rank dynamics in optimizing quantized LLMs. |
Copied to clipboard
| Challenge: | Large language models such as GPT-4 have limited their deployment in clinical settings . a novel framework for adapting SLMs into high-performing clinical models is needed . |
| Approach: | They propose a framework for adapting large language models into high-performing clinical models . they pre-instruct experts on relevant medical and clinical corpora and model merging . |
| Outcome: | The proposed framework outperforms the existing model on the CLUE+ benchmark on medical entities and radiology reports. |
Copied to clipboard
| Challenge: | Effectively resolving phonological ambiguities is crucial for robust natural language processing, as these ambiguity are pervasive in tasks ranging from speech-to-text, spelling correction, to offensive language detection. |
| Approach: | They propose a framework to enhance LLMs’ phonological capability through a multiple-stage training approach. |
| Outcome: | The proposed framework enables the base model to achieve comparable performance to a much larger model. |
Copied to clipboard
| Challenge: | Speculative decoding is a prominent technique for accelerating LLM inference by leveraging an auxiliary draft model, but its effectiveness is limited by the autoregressive nature of draft generation. |
| Approach: | They propose a method that integrates speculative draft generation directly within the target model using multi-stream attention. |
| Outcome: | The proposed method improves acceptance but also latency and speculation latency, limiting overall speedup. |
Copied to clipboard
| Challenge: | Existing RL methods suffer from reliability bottlenecks due to reward sparsity and intractable computations . d-TreeRPO provides fine-grained and verifiable step-wise reward signals . |
| Approach: | They propose a reliable reinforcement learning framework for diffusion large language models that leverages tree-structured rollouts and bottom-up advantage computation based on verifiable outcome rewards. |
| Outcome: | The proposed framework outperforms baseline models and achieves significant improvements across reasoning benchmarks. |
Copied to clipboard
| Challenge: | Large language models have shown remarkable capabilities, particularly in English, but for less prevalent languages, performance can be significantly lower, making additional adaptation paramount. |
| Approach: | They propose a new adaptation method based on iteratively merging multiple models fine-tuned on a subset of available training data that reduces forgetting while maintaining learning on the target domain. |
| Outcome: | The proposed method outperforms LLAMA-3-8B-based models in German and German while maintaining learning on the target domain. |
Copied to clipboard
| Challenge: | Existing instruction-following diffusion models are predominantly trained using an autoregressive paradigm. |
| Approach: | They propose a general instruction-following diffusion language model that outperforms contemporary instruction-tuned diffusion models and matches and sometimes exceeds strong autoregressive (AR) models. |
| Outcome: | The proposed model outperforms and sometimes exceeds existing autoregressive (AR) models on a number of tasks. |
Copied to clipboard
| Challenge: | Recent studies address safety-constrained online and offline preferences optimizations, but offline methods perform poorly in adaptively balancing safety and helpfulness. |
| Approach: | They propose a mixture of experts framework for safety-helpfulness dual Preference Optimization . they combine a single-preference enhanced direct preference optimization approach with a dynamic routing mechanism . |
| Outcome: | The proposed framework outperforms state-of-the-art methods in safety and helpfulness. |
Copied to clipboard
| Challenge: | EmoOmni is a data paradigm for omni-modal large language models that can be used for emotion reasoning. |
| Approach: | They propose a data paradigm that interleaves guided tokens into reasoning traces to enforce structured evidence extraction. |
| Outcome: | The proposed paradigm over-relys on a dominant modality while neglecting complementary cues. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have yielded impressive gains on mathematical reasoning benchmarks via supervised fine-tuning (SFT). |
| Approach: | They investigate the mechanisms behind SFT improvements in small-scale large language models by examining four key questions: (1) Are performance gains primarily due to format alignment rather than reasoning? (2) Can high-quality supervision encourage genuine reasoning? (4) Are format alignment gains consistent across model sizes and architectures? |
| Outcome: | The proposed models outperform the proprietary models on OlympiadBench and Omni-Math, but lack the brittleness of the models under perturbations to test their reasoning abilities. |
Copied to clipboard
| Challenge: | Recent fine-tuning approaches for large language models require supervised finetun on diverse datasets and follow different distributions. |
| Approach: | They propose a distribution edited model that integrates models individually trained on each data source with the base model using basic element-wise vector operations. |
| Outcome: | The proposed model outperforms baseline models on a variety of benchmarks and is cheaper than standard data mixing methods. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly deployed with task-specific adapters catering to multiple downstream applications. |
| Approach: | They propose a low-latency fused low-rank adapter that introduces zero latency overhead on top of the base model. |
| Outcome: | The proposed adapter reduces the inference time of the model by 2.5x . the proposed adapters are tested on 18 different tasks on different platforms . |
Copied to clipboard
| Challenge: | Existing approaches to enhance robustness of deep neural networks focus on perturbation . weak robustness is a problem for many types of adversarial attacks, authors say . |
| Approach: | They propose a lightweight framework for enhancing robustness by perturbing parameters of a model and diversifying adversarial example distributions among different models. |
| Outcome: | The proposed method can improve robustness against adversarial attacks while maintaining accuracy on clean data. |
Copied to clipboard
| Challenge: | Existing decoding-time defense methods suffer from limited generalization, high computational overhead, or significant utility degradation. |
| Approach: | They propose a decoding-time defense framework that leverages a pair of small contrastive models to estimate token-level safety signals by measuring divergence in their output distributions. |
| Outcome: | The proposed framework achieves near-zero attack success rates against a wide spectrum of advanced jailbreak attacks while maintaining the model’s helpfulness with minimal degradation. |
Copied to clipboard
| Challenge: | Modern language models rely on Reinforcement Learning from Human Feedback (RLHF) to encourage safe behaviors, but they remain vulnerable to adversarial attacks due to three key limitations: (1) the inefficiency and high cost of human annotation; (2) the vast diversity of potential adversarials; and (3) the risk of feedback bias and reward hacking. |
| Approach: | They propose an iterative adversarial training method that incorporates three key innovations to address these challenges. |
| Outcome: | Experiments on Mistral-7B-Instruct-v0.3 show that the proposed method significantly enhances robustness and reduces harmful outputs from 5.88% to 0.43%. |
Copied to clipboard
| Challenge: | Representation Fine-tuning (ReFT) is a proposed method for improving parameter efficiency . however, it yields suboptimal performance, as fixed-position representations have uncertain impact on outputs . |
| Approach: | They propose a method that fine-tunes critical representations in a low-rank linear subspace while freezing the base model. |
| Outcome: | The proposed method improves accuracy of LLaMA-2-7B and ReFT by 18.2 and 3.8 on GSM8K. |
Copied to clipboard
| Challenge: | Modern document retrieval embedding methods typically encode passages (chunks) from documents independently, often overlooking contextual information from the rest of the document. |
| Approach: | They propose a benchmark to evaluate retrieval models' ability to leverage document-wide context. |
| Outcome: | The proposed method significantly improves retrieval quality on ConTEB without sacrificing base model performance. |
Copied to clipboard
| Challenge: | Existing methods for long-horizon agents introduce the external memory module and look up the relevant information from the stored memory, which prevents the model from proactively managing its memory content and aligning with the agent’s overarching task objectives. |
| Approach: | They propose an algorithm which enables agents to autonomously manage their memory during interaction with environment and selectively retain crucial information. |
| Outcome: | Extensive experiments show that the proposed algorithm achieves absolute F1 score gains of 25.98 over the base model and 7.1 over the previous SOTA baseline while preserving task performance. |
Copied to clipboard
| Challenge: | a framework for model merging is proposed without additional training . task vectors from fine-tuned models exhibit a limited number of dominant singular values . |
| Approach: | They propose a framework for model merging based on low-rank estimation of task vectors without access to the base model. |
| Outcome: | The proposed framework improves models without additional training without additional inputs. |
Copied to clipboard
| Challenge: | Existing approaches to improve latency via skipping layers have limitations . fiRST is a model-agnostic framework that reduces inference latency while maintaining quality . |
| Approach: | They propose a model-agnostic framework that skips transformer layers during decoding . it is fully compatible with KV caching, enabling faster decoding while maintaining quality . |
| Outcome: | a new framework reduces inference latency by using layer-specific routers to skip transformer layers during decoding. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning language models are efficient when adapting to a single dataset. |
| Approach: | They propose to use an ensemble method for fine-tuning a language model to multiple datasets instead of a single adapter per task. |
| Outcome: | The proposed method improves performance on multiple datasets while preserving low-rank adaptation properties. |
Copied to clipboard
| Challenge: | Results show that supervised fine-tuning and preference finetunation are the most efficient approaches for large language models. |
| Approach: | They propose to use Supervised Finetuning and Preference Finetunes to optimize training data budgets for Large Language Models. |
| Outcome: | The proposed approach improves performance on math tasks by 15% on the most expensive model, 1,000 examples. |
Copied to clipboard
| Challenge: | Existing approaches focus on retrieval augmentation and focus on the quality of the output . Existing methods focus on generating a highly specific declarative statement ignoring the underlying reasoning process behind ideation. |
| Approach: | They propose a large language model that generates evidence-based hypotheses using literature-guided reasoning and a multi-task setting. |
| Outcome: | The proposed model outperforms the base model and generates evidence-grounded hypotheses with high feasibility and impact as judged by human experts. |
Copied to clipboard
| Challenge: | Reinforcement learning is emerging as a primary driver for improving language model reasoning capabilities. |
| Approach: | They propose a method for explicitly up-weighting rare but correct solutions to overcome rank bias in group relative policy optimization (GRPO) . |
| Outcome: | The proposed method mitigates rank bias and improves pass@N across a large range of N in both synthetic and real theorem proving settings. |
Copied to clipboard
| Challenge: | Standard language models employ unique, monolithic embeddings for each token, limiting their ability to capture multifaceted meanings. |
| Approach: | They propose a compositional structure that accumulates diverse semantic facets for tokens . they apply this representational scheme to standard transformer architectures and a biomedical domain benchmark . |
| Outcome: | The proposed representational scheme achieves extreme compression in embedding parameters while maintaining >95% task performance relative to the base model. |
Copied to clipboard
| Challenge: | Model merging has emerged as a promising technique for combining fine-tuned models into a single expert model without retraining. |
| Approach: | They propose a model merging technique that preserves weak model knowledge . they define mergeability as a property of model updates that captures how well they retain trained knowledge when merged with other model updates. |
| Outcome: | The proposed method preserves weak knowledge in the base model. |
Copied to clipboard
| Challenge: | Existing methods that prune or employ early stopping to reduce latency often compromise reasoning reliability. |
| Approach: | They propose a shortcut decoding framework that integrates probes over internal hidden states with step-level entropy to detect convergence of reasoning during generation and adaptively selects between a fast-exit path and a stability-verified path to remove redundant steps while preserving answer correctness. |
| Outcome: | The proposed framework reduces token usage by approximately 35% and maintains accuracy comparable to full CoT decoding. |
Copied to clipboard
| Challenge: | Existing studies focus on building text-only agents in synthetic environments where the reward signals are clearly defined. |
| Approach: | They propose a multimodal web agent that can autonomously conduct real-world exploration and improve itself after each iteration. |
| Outcome: | The proposed agent improves itself after each iteration, demonstrating strong performance across multiple test sets. |
Copied to clipboard
| Challenge: | Existing safety defenses typically intervene internally within the generative model, but suffer from severe concept entanglement, leading to degradation of benign generation quality. |
| Approach: | They propose a structurally isolated safety module that performs external, interpretable rectification without modifying the base model. |
| Outcome: | The proposed module performs external, interpretable rectification without modifying the base model. |
Copied to clipboard
| Challenge: | Reinforcement learning (RL) has emerged as a powerful paradigm for improving the reasoning capabilities of large language models. |
| Approach: | They propose a pipeline that automatically discovers thinking token patterns with reasoning primitives and curates SFT datasets to prepare LLMs for RL. |
| Outcome: | The proposed pipeline outperforms baseline methods on mathematical and logical reasoning benchmarks on RL tasks. |
Copied to clipboard
| Challenge: | Existing methods for label detection and explanation generation have been limited in understanding complex issues . identifying propaganda and hate in memes is essential for combating misinformation and minimizing harm . |
| Approach: | They propose an explanation-enhanced dataset for propaganda memes in Arabic and hateful memes on English to solve these tasks. |
| Outcome: | The proposed model outperforms the current state-of-the-art in label detection and explanation generation. |
Copied to clipboard
| Challenge: | AutoSDT-5K is the only automatically collected and the largest open dataset for data-driven scientific discovery. |
| Approach: | They propose an automatic pipeline that collects high-quality coding tasks in real-world data-driven discovery workflows. |
| Outcome: | The proposed pipeline synthesizes accurate tasks and tasks from a dataset of 5,404 tasks covering four scientific disciplines and 756 Python packages. |
Copied to clipboard
| Challenge: | Low-Rank Adaptation (LoRA) is a new approach to fine-tuning large language models . adapters are lightweight, task specific modules that can be used for adapters in latency-sensitive settings. |
| Approach: | They propose a low-rank adapter with a weight sharing mechanism that reduces latency by 40% . they analyze LoRA adapters on GPUs and identify segmented function calls as the primary source of latency. |
| Outcome: | The proposed adapter reduces latency to about 40% of the gap between the unmerged LoRA and the base model while maintaining parameter efficiency and comparable accuracy. |
Copied to clipboard
| Challenge: | prevailing methods rely on hand-crafted or pre-specified strategies and struggle to balance efficiency, imperceptibility, and security, particularly at high embedding rates. |
| Approach: | They propose an agent-driven self-evolving framework that is the first to realize self-changing steganographic strategies by automatically discovering, composing, and adapting strategies at inference time. |
| Outcome: | The proposed framework achieves 42.2% perplexity and 1.6% anti-steganalysis performance over SOTA methods at high embedding rates. |
Copied to clipboard
| Challenge: | Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning . however, the recipe introduces a significant risk of capability regression, where models forget foundational skills after prolonged training without employing regularization strategies. |
| Approach: | They propose a replay strategy with dynamic objective reweighting for general knowledge preservation using short-horizon signals of convergence and instability. |
| Outcome: | The proposed method preserves general capabilities and improves reasoning . it can be applied to existing RLVR pipelines without training additional models or tuning . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have made remarkable progress through Reinforcement Learning with Verifiable Rewards (RLVR) however, external supervision remains a bottleneck for tasks and domains for which supervised data are scarce or non-existent. |
| Approach: | They propose a novel dual-play framework that adversarially trains two models initialized from the same base model. |
| Outcome: | The proposed framework improves the math reasoning performance of large language models. |
Copied to clipboard
| Challenge: | Recent Large Audio Language Models (LALMs) have shown strong capabilities in audio understanding, yet their reasoning remains vulnerable to perceptual errors. |
| Approach: | They propose a large-scale dataset for **Perception-Aware Question Answering** that uses a hierarchical decoupling strategy to separate speech from environmental sounds and distinguishes among multiple speakers. |
| Outcome: | The proposed model improves on MMAU-mini, MMAR, and PAQA while maintaining comparable performance on multiple benchmarks. |
Copied to clipboard
| Challenge: | Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models beyond the training domain, while supervised fine-tuning (SFT) frequently leads to general capabilities forgetting. |
| Approach: | They propose a feature-level mechanistic analysis methodology to probe RL generalization using a controlled experimental setup. |
| Outcome: | The proposed method identifies a compact, task-agnostic set of features that directly mediate generalization across diverse tasks. |
Copied to clipboard
| Challenge: | Recent reasoning models fail to capture structural constraints in complex settings. |
| Approach: | They propose a visual-based reasoning system that integrates executable visual construction into multi-turn reasoning via end-to-end reinforcement learning. |
| Outcome: | The proposed model outperforms strong text-only chain-of-thought models on seven mathematical benchmarks and improves by 13.12% on AIME 2025 and 11.00% on BeyondAIME. |
Copied to clipboard
| Challenge: | Recent advances in large language models have driven reasoning performance . low-resource distillation can boost models' performance, but a framework is missing . |
| Approach: | They conduct a controlled experiment to find out why low-resource distillation can boost model performance . they find that distillation enhances the presence of advanced cognitive behaviors . |
| Outcome: | The proposed model shows more flexible reasoning, the authors show . they show that distillation enhances the presence of advanced cognitive behaviors . |
Copied to clipboard
| Challenge: | Existing benchmarks lack systematically paired instances across modalities, making it difficult to compare genuine arithmetic limits . a model that computes 4736 may fail on a nearby instance like 8967, despite a well-tuned internal router. |
| Approach: | They propose a controlled multimodal multiplication benchmark that factorially varies digit length, digit sparsity, representation, and modality with paired instances from a reproducible generator. |
| Outcome: | The proposed model can perceive numerical content across modalities but fails to perform exact multi-digit multiplication when presented as numerals, number words, images, or in audio form. |
Copied to clipboard
| Challenge: | Recent discussions suggest that further progress will come from scaling the right structure, not merely parameters or data, while preserving acquired knowledge. |
| Approach: | They propose a width upscaling architecture that inserts lightweight expansions into linear modules while freezing all pre-trained parameters. |
| Outcome: | The proposed architecture reduces severe forgetting while learning new knowledge on a controlled synthetic biography benchmark. |
Copied to clipboard
| Challenge: | Existing methods for text-to-image alignment evaluation rely on coarse-grained metrics or static Question Answering pipelines that lack fine-grounded interpretability and struggle to reflect human preferences. |
| Approach: | They propose a reinforcement-guided visual reasoning framework for element-level text-to-image alignment evaluation. |
| Outcome: | The proposed framework achieves state-of-the-art results on four benchmarks and surpasses the strong proprietary Gemini 3 Pro and Training-based baselines. |